Papers with speech recognition

83 papers
Cross-Lingual Transfer Learning for Speech Translation (2025.naacl-short)

Copied to clipboard

Challenge: Increasing interest in building multilingual foundation models for NLP and speech research has led to limited data collection for training ST systems.
Approach: They propose to use Whisper to explore the behavior of multilingual speech foundation models with restricted data.
Outcome: The proposed model can translate to Chinese with a single language, and it can perform transcriptions in other languages.
Simultaneous Translation (2020.emnlp-tutorials)

Copied to clipboard

Challenge: Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation.
Approach: This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation.
Outcome: This tutorial will examine the design and evaluation of policies for simultaneous translation .
Deep Bayesian Natural Language Processing (P19-4)

Copied to clipboard

Challenge: Introduction to deep Bayesian learning for natural language addresses the fundamentals of statistical models and neural networks.
Approach: This tutorial addresses the advances in deep Bayesian learning for natural language . it focuses on advanced Bayessian models and deep models . authors present case studies and domain applications to tackle different issues .
Outcome: This tutorial focuses on advanced Bayesian models and deep models for natural language . case studies and domain applications are presented to tackle different issues in deep Bayessian processing, learning and understanding.
An adaptable task-oriented dialog system for stand-alone embedded devices (P19-3)

Copied to clipboard

Challenge: a proposed speech-based task-oriented dialogue system is built on a small embedded device . the system does not require internet connectivity because all components run locally on the device - a cost-effective solution .
Approach: They propose a spoken-language end-to-end task-oriented dialogue system for small embedded devices such as home appliances.
Outcome: The proposed system is based on a demo run offline on swiss raspberry pi . it eliminates privacy risks and eliminates server costs and latency .
From dictations to clinical reports using machine translation (N18-3)

Copied to clipboard

Challenge: Medical dictation is one of the most common ways to document clinical encounters.
Approach: They propose a machine callytranslation technique that automates post-processing tasks . they show that it outperforms conventional systems in correcting errors .
Outcome: The proposed method outperforms conventional systems in many tasks while being much simpler to maintain.
KT-Speech-Crawler: Automatic Dataset Construction for Speech Recognition from YouTube Videos (D18-2)

Copied to clipboard

Challenge: KT-Speech-Crawler is an automated dataset building tool for speech recognition.
Approach: They propose an approach for automatic dataset construction for speech recognition by crawling YouTube videos.
Outcome: The proposed algorithm can obtain 150 hours of transcribed speech in a day with an estimated 3.5% word error rate.
FPI: Failure Point Isolation in Large-scale Conversational Assistants (2022.naacl-industry)

Copied to clipboard

Challenge: Large-scale conversational assistants can cause errors in their modules . a machine learning system can analyze large volumes of data and isolate the source of error .
Approach: They propose a machine learning system that embeds incoming request and context using pre-trained transformer models and encodes additional metadata features to output failure point predictions.
Outcome: The proposed system obtains 92.2% of human performance while scaling to analyze the entire traffic in 8 different languages of a large-scale conversational assistant.
RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition (P18-4)

Copied to clipboard

Challenge: Using RETURNN, we train and decode attention models for translation and speech recognition.
Approach: They propose a layer-wise pretraining scheme for recurrent attention models and show its significant effect on deep recurrence encoder networks.
Outcome: The proposed training and decoding scheme improves 1% on expected training and improves on WMT 2017 and Switchboard.
Practical Application of Domain Dependent Confidence Measurement for Spoken Language Understanding Systems (N18-3)

Copied to clipboard

Challenge: a confidence score is a scalar quantity that measures the reliability of an automatic system.
Approach: They propose to use a confidence measure to evaluate the reliability of an SLU system . they build confidence models for three different types of dialogue states .
Outcome: The proposed model can be used to reject low-confidence SLU results in real-world scenarios.
A Research Platform for Multi-Robot Dialogue with Humans (N19-4)

Copied to clipboard

Challenge: a new research platform supports spoken dialogue interaction with multiple robots . a ground robot and an aerial robot are used to perform search and rescue tasks .
Approach: They propose a platform that supports spoken dialogue interaction with multiple robots . they use existing tools for speech recognition and dialogue management .
Outcome: The proposed platform supports spoken dialogue interaction with multiple robots in a search and rescue scenario.
Neural Text Normalization with Subword Units (N19-2)

Copied to clipboard

Challenge: Text normalization (TN) is an important step in conversational systems.
Approach: They frame text normalization as a machine translation task and tackle it with sequence-to-sequence models.
Outcome: The proposed model normalizes written text to its spoken form to facilitate speech recognition and text-to-speech synthesis.
PyOpenDial: A Python-based Domain-Independent Toolkit for Developing Spoken Dialogue Systems with Probabilistic Rules (D19-3)

Copied to clipboard

Challenge: a recent development of spoken dialogue systems has enabled deep learning to achieve state-of-the-art performance.
Approach: They propose a Python-based domain-independent, open-source toolkit for spoken dialogue systems.
Outcome: The proposed toolkit extends OpenDial's Java-based architecture and provides new functions for neural dialogue state tracking and action planning.
Simultaneous Speech-to-Text Translation Web Application for Estonian (2026.eacl-demo)

Copied to clipboard

Challenge: a new open-source web application for simultaneous speech-to-text translation is developed for Estonian . the system translates live Estonian speech into English, Russian, and Ukrainian text, and also supports English-to Estonian translation.
Approach: They propose a web application that combines streaming speech recognition with a simultaneous translation model.
Outcome: The proposed system outperforms existing streaming speech recognition systems in Estonian-to-English translation.
TLT-school: a Corpus of Non Native Children Speech (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of speech utterances collected in schools of northern italy is being used to assess the performance of students learning both English and German.
Approach: a corpus of speech utterances collected in schools of northern italy is described . the corpus is going to be freely distributed to scientific community .
Outcome: The corpus of speech utterances collected in schools of northern italy is a "Trentino Language Testing" in schools" the data are used to assess the performance of students learning English and German .
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training (2024.findings-naacl)

Copied to clipboard

Challenge: Recent advances in language modeling have led to the emergence of large language models capable ofvarious natural language processing tasks.
Approach: They propose a multi-instructional training approach that integrates a large language model with a speech encoder to harness the capabilities of LLMs for speech recognition and beyond.
Outcome: The proposed model can be trained and aligned with a multilingual LLM on 1900 hours of transcribed data from 139 languages.
On the Use of External Data for Spoken Named Entity Recognition (2022.naacl-main)

Copied to clipboard

Challenge: Named entity recognition (NER) tasks require large labeled datasets to perform . compared to prior work, relative improvements in F1 of up to 16% are found .
Approach: They propose to use self-training, knowledge distillation, and transfer learning to learn SLU models . they compare pipeline and pipeline approaches to find out how to use external data .
Outcome: The proposed models improve performance beyond pre-trained models in resource-constrained settings . the best baseline model is a pipeline approach, while the best performance is achieved by an E2E model.
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs (2025.acl-short)

Copied to clipboard

Challenge: Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech.
Approach: They propose an EMG adaptor module that maps EMG features to an LLM's input space and achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task.
Outcome: The proposed module achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task.
A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition Baseline (2021.eacl-main)

Copied to clipboard

Challenge: The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders.
Approach: They propose to build an open-source Kazakh speech corpus for the Kazakh language that contains over 153,000 transcribed audio . they describe the data collection and preprocessing procedures followed by a description of the database specifications.
Outcome: The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders.
InTriage: Intelligent Telephone Triage in Pre-Hospital Emergency Care (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing TT processes face challenges such as incomplete data collection, communication barriers, and manual errors, leading to high over-triage and under-triages rates.
Approach: They propose to use an AI-driven multilingual TT system to provide decision support for triage.
Outcome: The proposed system achieves word error rate of 14.57% for speech recognition and an F1 score of 73.34% for key information extraction.
Efficient Sequence Learning with Group Recurrent Networks (N18-1)

Copied to clipboard

Challenge: Recurrent neural networks have achieved state-of-the-art results in many artificial intelligence tasks, such as language modeling, neural machine translation and speech recognition.
Approach: They propose an efficient architecture to improve the efficiency of such RNN model training by adopting the group strategy for recurrent layers while exploiting the representation rearrangement strategy between layers as well as time steps.
Outcome: The proposed architecture achieves comparable or better accuracy compared with baselines, with a much smaller number of parameters and at a lower computational cost.
Pisets: A Robust Speech Recognition System for Lectures and Interviews (2025.naacl-industry)

Copied to clipboard

Challenge: Sustainable speech recognition systems are essential for scientists, journalists, and anyone processing audio recordings of interviews and meetings.
Approach: They propose a speech-to-text system "Pisets" which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model.
Outcome: The proposed system ensures robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model.
Learning Hidden Unit Contribution for Adapting Neural Machine Translation Models (N18-2)

Copied to clipboard

Challenge: In this paper we explore the use of Learning Hidden Unit Contribution for neural machine translation.
Approach: They propose to use Learning Hidden Unit Contribution for the task of neural machine translation.
Outcome: The proposed method achieves improvements of up to 2.6 BLEU points over a general system . it also achieves up to 6 BLUE points if the initial system has been trained on out-of-domain data .
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation (2021.acl-long)

Copied to clipboard

Challenge: VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST.
Approach: They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages.
Outcome: The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours.
Multimodal fusion via cortical network inspired losses (2022.acl-long)

Copied to clipboard

Challenge: Recent work in deep fusion models has led to substantial improvements over unimodal approaches in areas like speech recognition, emotion recognition and analysis.
Approach: They propose to introduce neural dependencies into the loss functions to allow for fusion of different modalities while keeping the model complexity manageable.
Outcome: Experiments on multimodal sentiment analysis tasks show that the proposed approach provides a consistent performance boost.
Unified Speech-Text Pre-training for Speech Translation and Recognition (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to pre-train speech and text use unlabeled data to learn universal feature representations.
Approach: They propose a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition.
Outcome: The proposed method achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task.
BIG-C: a Multimodal Multi-Purpose Dataset for Bemba (2023.acl-long)

Copied to clipboard

Challenge: Bemba is the most populous language of Zambia but lacks resources for research . despite its significance, Bemba remains under-resourced and lacking in high-quality data and resources for NLP experiments and language technologies.
Approach: They propose a large multimodal dataset for Bemba that includes images, transcriptions and translations.
Outcome: The proposed dataset is based on images, transcriptions and translations of Bemba speakers . it provides baselines on speech recognition, machine translation and speech translation tasks .
Self-Attentional Models for Lattice Inputs (P19-1)

Copied to clipboard

Challenge: Existing work has extended recurrent neural networks to model lattice inputs but these models suffer from slow computation speeds.
Approach: They propose to extend the paradigm of self-attention to handle lattice inputs by adding probabilistic reachability masks that incorporate latticae structure into the model and support lattics if available.
Outcome: The proposed model outperforms baseline models while being much faster to compute than previous models.
AfriVox: Probing Multilingual and Accent Robustness of Speech LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Recent advances in multimodal and speech-native large language models have delivered impressive speech recognition, translation, understanding, and question-answering capabilities for high-resource languages.
Approach: They propose to benchmark African languages and African-accented French, Arabic, and 100+ African English accents across 20 African languages.
Outcome: The proposed model outperforms traditional speech transcription and translation models in African languages and non-native French or English accents.
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents (2024.findings-eacl)

Copied to clipboard

Challenge: AccentFold uses spatial relationships to improve speech recognition for accented speech . existing methods for accent recognition have been limited due to data scarcity and budget constraints .
Approach: They propose a method that exploits spatial relationships between learned accent embeddings to improve downstream automatic speech recognition.
Outcome: The proposed method outperforms baseline methods in accented speech training.
Large Margin Neural Language Model (D18-1)

Copied to clipboard

Challenge: Conventionally, neural language models are trained by minimizing perplexity (PPL) on grammatical sentences.
Approach: They propose a large margin criterion for training neural language models by minimizing perplexity on grammatical sentences and propose enlarged margins for task-specific training.
Outcome: The proposed method gains up to 1.1 WER reduction for speech recognition and 1.0 BLEU increase for machine translation.
Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization (P19-1)

Copied to clipboard

Challenge: Existing methods to regularize multimodal data are imperfect due to imperfect modalities, missing entries or noise corruption.
Approach: They propose a method to regularize multimodal data by tensor rank minimization . they use correlations between time and modalities to generate low-rank tenses .
Outcome: The proposed model achieves good results across various levels of imperfection.
AdaTranS: Adapting with Boundary-based Shrinking for End-to-End Speech Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: End-to-end speech translation (ST) models need large amount of training data to perform well.
Approach: They propose a shrinking mechanism to mitigate the length mismatch between speech and text features by predicting word boundaries.
Outcome: The proposed method achieves better performance on the MUST-C dataset, with higher inference speed and lower memory usage.
Automatic Partitioning of a Code-Switched Speech Corpus Using Mixed-Integer Programming (2024.lrec-main)

Copied to clipboard

Challenge: Currently, partitioning speech corpora is done by hand, but this is not feasible for the dataset under investigation.
Approach: They propose to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions using mixed-integer linear programming.
Outcome: The proposed method allows to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions while maintaining a fixed number of speakers and a specific amount of codeswitching speech in the development and test partitions.
When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP (2024.acl-long)

Copied to clipboard

Challenge: despite its crucial role in research experiments, code correctness is often presumed on the perceived quality of results.
Approach: They propose to promote code-quality checklists to promote coding best practices . they propose to fix bugs in conformer implementations to mitigate this risk .
Outcome: The proposed checklists aim to promote coding best practices and improve software quality within the NLP community.
Improved Speech Representations with Multi-Target Autoregressive Predictive Coding (2020.acl-main)

Copied to clipboard

Challenge: Autoregressive coding targets are used to learn meaningful representations from unlabeled speech.
Approach: They propose a method that trains an autoregressive RNN to generate an unseen future frame given a context such as recent past frames.
Outcome: The proposed method can learn representations from unlabeled speech.
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution (2025.findings-acl)

Copied to clipboard

Challenge: x-vector (speaker recognition PTM) achieves the highest performance in prosodic tasks . despite its low parameter, x vector captures unique prosodic characteristics of the sources .
Approach: They propose to use SOTA speech pre-trained models to capture prosodic sig-natures of generative sources for audio deepfake source attribution.
Outcome: The proposed model captures prosodic sig-natures of generative sources better than other models on ASVSpoof and CFAD.
Indigenous language technologies in Canada: Assessment, challenges, and successes (C18-1)

Copied to clipboard

Challenge: There are approximately 60 Indigenous languages currently spoken in Canada.
Approach: They examine which technologies have been developed and which are feasible to develop for the 60 Indigenous languages spoken in Canada.
Outcome: The proposed technologies are based on the existing technologies and are feasible for most or all of these languages.
Fine-Grained Grounding for Multimodal Speech Recognition (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models rely on global visual features that represent the entire image, but localizing the relevant regions of the image will make it possible to recover a larger set of words, such as adjectives and verbs.
Approach: They propose a multimodal automatic speech recognition system that uses visual information from different parts of the image to ground the speech in the visual context.
Outcome: The proposed model improves over approaches that use global visual features and localizes the correct proposals.
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself.
Approach: They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations.
Outcome: The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform.
Composing Finite State Transducers on GPUs (P18-1)

Copied to clipboard

Challenge: Weighted finite state transducers (FSTs) are used in language processing . a GPU implementation of the composition operation is currently under development .
Approach: They propose a GPU implementation of the composition operation for weighted finite state transducers.
Outcome: The proposed approach achieves speedups of up to 6 times over the serial implementation and 4.5 times over OpenFST on the GPU.
Investigating the Emergent Audio Classification Ability of ASR Foundation Models (2024.naacl-long)

Copied to clipboard

Challenge: Text and vision foundation models can perform many tasks in a zero-shot setting . however, there has been little work on the zero-shoot abilities of ASR foundation models .
Approach: They investigate the ability of ASR foundation models to perform zero-shot audio classification using text prompts and a decoding probability generator.
Outcome: The proposed model outperforms state-of-the-art models on audio classification datasets without training them on extra data or adding any parameters.
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to metric meta-evaluation focus on general statements about absolute and relative quality of metrics across arbitrary system outputs, but in practice, metrics are applied in highly contextual settings.
Approach: They propose a method for contextual metric meta-evaluation by comparing local metric accuracy.
Outcome: The proposed method compares the local metric accuracy of evaluation metrics across translation, speech recognition, and ranking tasks.
Fluent Translations from Disfluent Speech in End-to-End Speech Translation (N19-1)

Copied to clipboard

Challenge: Disfluency removal is an intermediate step between speech recognition and machine translation (MT) with the rise of end-to-end speech translation systems, disfluency recognition and removal needs to be incorporated into the model architectures or handled as a post-processing step.
Approach: They propose to use a sequence-to-sequence model to translate from noisy, disfluent speech to fluent text with disfluencies removed using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset.
Outcome: The proposed model generates fluent translations from disfluent speech using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset.
Best of Both Worlds: Making High Accuracy Non-incremental Transformer-based Disfluency Detection Incremental (2021.acl-long)

Copied to clipboard

Challenge: Currently, Transformer-based text classifiers are not suitable for live incremental processing, operating only on the level of complete sentence inputs.
Approach: They propose to introduce a method for word-by-word left-to-right incremental processing to Transformers such as BERT, models without an intrinsic sense of linear order.
Outcome: The proposed method maintains high non-incremental performance while operating strictly incrementally.
DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers (2024.findings-naacl)

Copied to clipboard

Challenge: Existing interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity.
Approach: They propose a method to analyze encoder-decoder Transformers by using the decoder module Model Output encoder to cross-attend representations of intermediate encoder activations instead of using the default output.
Outcome: The proposed method maps uninterpretable representations to human-interpreted sequences of words or symbols, shedding new light on the information flow in this popular but understudied class of models.
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models.
Approach: They propose a single-stage joint speech-text SFT approach for training SpeechLMs . their model combines text-only SFT data with three types of speech-related data .
Outcome: The proposed model outperforms previous SpeechLMs on speech-based QA tasks while maintaining original speech-only capabilities.
Curriculum Pre-training for End-to-End Speech Translation (2020.acl-main)

Copied to clipboard

Challenge: End-to-end speech translation requires a powerful encoder to transcribe, understand and learn cross-lingual semantics simultaneously.
Approach: They propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages.
Outcome: The proposed method improves on En-De and En-Fr speech translation benchmarks.
Meta-Transfer Learning for Code-Switched Speech Recognition (2020.acl-main)

Copied to clipboard

Challenge: Increasing number of people in the world today speak a mixed-language as a result of being multilingual.
Approach: They propose a method to transfer learn on a code-switched speech recognition system by extracting information from high-resource monolingual datasets.
Outcome: The proposed model outperforms baselines on speech recognition and language modeling tasks and is faster to converge.
A Resource for Computational Experiments on Mapudungun (2020.lrec-1)

Copied to clipboard

Challenge: Low-resource languages still lag behind in documenting endangered languages . a large corpus of culturally significant conversations is available for computational experiments .
Approach: They propose a resource for computational experiments on Mapudungun, a polysynthetic indigenous language spoken in Chile.
Outcome: The proposed corpus provides 142 hours of culturally significant conversations in Mapudungun . the language is spoken by the Mapuche people of southern Chile and western argentina .
Fashioning Local Designs from Generic Speech Technologies in an Australian Aboriginal Community (2022.coling-1)

Copied to clipboard

Challenge: Recent research has focused on low-resource languages and the transcription bottleneck paradigm.
Approach: They propose to use a spoken term detection system to train a speech recognition system in an Aboriginal community to reach better comprehension and engagement from Aboriginal participants.
Outcome: The proposed system can be implemented in an Aboriginal community and reach better comprehension and engagement from Aboriginal participants.
Language Technology Programme for Icelandic 2019-2023 (2020.lrec-1)

Copied to clipboard

Challenge: a new national language technology programme for Icelandic is described . the programme aims to make Icelandic usable in communication and interactions in the digital world .
Approach: They describe a new national language technology programme for Icelandic . the programme aims to make Icelandic usable in communication and interactions in the digital world .
Outcome: The proposed programme aims to make Icelandic usable in communication and interactions in the digital world.
Evaluation of African American Language Bias in Natural Language Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that large language generation models disadvantaging African American Language (AAL) can be biased for certain language varieties, but there is little research on the impact of these biases on other languages.
Approach: They evaluate how well LLMs understand African American Language (AAL) in comparison to white Mainstream English (WME) using a dataset of AAL texts from a variety of regions and contexts, they find dialectal bias in six pre-trained LLM.
Outcome: The proposed models understand African American language in comparison to white mainstream English (WME) the proposed models have performance gaps on two tasks that are not matched by the model.
Learning to Detect Noisy Labels Using Model-Based Features (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to reduce label noise rely on heuristics and sample losses.
Approach: They propose a method that transfers the noise distribution to a clean set and trains a model to distinguish noisy labels from clean ones using model-based features.
Outcome: Empirically, the proposed approach improves over strong baselines on a wide range of tasks including text classification and speech recognition.
LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data .
Approach: They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books.
Outcome: The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation.
Preparing Data from Psychotherapy for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: mental health care is a demanding occupation, resulting in a severe gap in patient-centered care . a recent study shows that natural language processing can extract certain aspects of human-human communication.
Approach: They propose to use data from psychotherapy sessions to help improve quality of care . they use feedback and cooperation annotations to assess quality of therapy sessions .
Outcome: The proposed method aims to analyse psychotherapy data and assess its quality . it aims at identifying what qualifies for good feedback or cooperation in therapy sessions .
SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding Tasks (2023.acl-long)

Copied to clipboard

Challenge: Spoken language understanding (SLU) tasks have received little attention and resources compared to lower-level tasks like speech and speaker recognition.
Approach: They propose annotated SLU benchmark tasks based on freely available speech data to complement existing benchmarks and address gaps in the evaluation landscape.
Outcome: The proposed benchmarks complement existing benchmarks and address gaps in the evaluation landscape.
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for speech recognition suffer from the synthetic-to-real gap . existing methods suffer from this distributional shift due to acoustic mismatches .
Approach: They propose to use task arithmetic to fine-tune an ASR model on synthetic data to mitigate the synthetic-to-real gap.
Outcome: The proposed method shows an improvement of 10.03% over baselines on the SLURP dataset.
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development.
Approach: They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments.
Outcome: The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages.
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: the Huqariq corpus is a multilingual collection of speech from native Peruvian languages . the project is designed to preserve endangered languages in the public domain .
Approach: They propose to use crowdsourcing to collect transcribed audio from native Peruvian languages . they propose to do 220 hours of speech recognition experiments to verify quality .
Outcome: The Huqariq corpus is a multilingual collection of speech from native Peruvian languages . the project is expected to reach 20 native languages out of 48 native languages by 2022 .
A Joint Approach to Compound Splitting and Idiomatic Compound Detection (2020.lrec-1)

Copied to clipboard

Challenge: Compounding is a common word-formation process in Germanic languages . high productivity and low corpus frequency of compounds increase vocabulary size .
Approach: They develop a deep learning-based approach to noun compound splitting and idiomatic compound detection for the German language.
Outcome: The proposed approach outperforms the current state of the art in noun compound splitting and idiomatic compound detection for the German language.
Where are we in Named Entity Recognition from Speech? (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition is usually made through a pipeline process that consists of processing audio and applying a NER to the audio outputs.
Approach: They propose an original 3-pass approach and explore the capability of an E2E system to do structured NER.
Outcome: The proposed system performs better than the current pipeline approach.
Multilingual and Cross-Lingual Intent Detection from Spoken Data (2021.emnlp-main)

Copied to clipboard

Challenge: a systematic study on multilingual and cross-lingual intent detection from spoken data is presented . current work on intent detection is limited to English, and standard benchmarks exist only in English.
Approach: They present a systematic study on multilingual and cross-lingual intent detection from spoken data.
Outcome: The proposed resource is called MInDS-14, and it provides strong intent detection in most target languages.
Collection and Analysis of Code-switch Egyptian Arabic-English Speech Corpus (L18-1)

Copied to clipboard

Challenge: despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech.
Approach: They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective.
Outcome: The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective.
Lattice Transformer for Speech Translation (P19-1)

Copied to clipboard

Challenge: Recent advances in sequence modeling have highlighted the strengths of the transformer architecture.
Approach: They propose a general lattice transformer for speech translation where the input is the output of the automatic speech recognition (ASR) they propose 'controllable' lattica attention mechanism to consume latent representations.
Outcome: The proposed model outperforms baseline and lattice LSTM on the Chinese-English translation task.
Speech Translation and the End-to-End Promise: Taking Stock of Where We Are (2020.acl-main)

Copied to clipboard

Challenge: Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach.
Approach: They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity .
Outcome: This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity.
Open Terminology Management and Sharing Toolkit for Federation of Terminology Databases (2022.lrec-1)

Copied to clipboard

Challenge: Terminology is also needed in AI applications such as machine translation, speech recognition, information extraction, and other natural language processing tools.
Approach: They propose a terminology management solution that facilitates standards-based sharing and management of terminology resources by providing the EuroTermBank Toolkit.
Outcome: The EuroTermBank Toolkit facilitates standards-based sharing and management of terminology resources by participating in federated databases.
wav2vec-S: Adapting Pre-trained Speech Models for Streaming (2024.findings-acl)

Copied to clipboard

Challenge: Pre-trained speech models have advanced speech-related tasks, including speech recognition and translation.
Approach: They propose a pre-trained speech model that incorporates modifications to ensure consistent speech representations during training and inference phases for streaming speech inputs.
Outcome: The proposed model outperforms baseline models on speech recognition and translation tasks and achieves a superior balance between quality and latency.
VAST: A Corpus of Video Annotation for Speech Technologies (L18-1)

Copied to clipboard

Challenge: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Approach: The Video Annotation for Speech Technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development .
Outcome: The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech detection, language identification, speaker identification, and speech recognition .
CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition (2022.lrec-1)

Copied to clipboard

Challenge: In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety.
Approach: They propose a dataset for in-car command recognition in the cantonese language with both video and audio data.
Outcome: The proposed model can achieve a considerable quality on the clean test set, but the speech recognition quality on noisy data is still inferior.
Evaluation of Off-the-shelf Speech Recognizers Across Diverse Dialogue Domains (2020.lrec-1)

Copied to clipboard

Challenge: a recent study evaluated off-the-shelf automatic speech recognition systems . current state-of-the art systems perform poorly in domains that require special vocabulary and language models .
Approach: They evaluate off-the-shelf automatic speech recognition systems across different dialogue domains . they use data collected from deployed spoken dialogue systems and human-human conversations .
Outcome: The evaluation is aimed at non-experts with limited experience in speech recognition . the results show that the performance of each speech recognizer can vary significantly depending on the domain .
Optimized Tokenization for Transcribed Error Correction (2023.emnlp-main)

Copied to clipboard

Challenge: transcribed-like data is often used to correct recurring errors, but training with synthetic data is difficult.
Approach: They propose to use synthetic transcribed-like data to train error correction models . they show that synthetic data outperforms the common approach of random perturbations .
Outcome: The proposed method outperforms the common method using random perturbations in transcribed data and language-specific adjustments to the vocabulary of a BPE tokenizer.
Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: In an aging society, a highly accurate speech recognition system is needed for use in electronic devices for the elderly but this cannot be achieved using conventional speech recognition systems due to the unique features of the speech of elderly people.
Approach: They construct a new corpus of elderly Japanese speech from existing Japanese speech corpora and train them using existing data.
Outcome: The proposed models achieve word error rates (WER) as low as 13.38%, exceeding the results of the previous study.
Is Spoken Hungarian Low-resource?: A Quantitative Survey of Hungarian Speech Data Sets (2024.lrec-main)

Copied to clipboard

Challenge: Existing data sets in Hungarian are limited in quality and quality . however, it is difficult to train a modern automatic speech recognition system with thousands of hours of transcribed speech.
Approach: They propose to analyze available speech data sets in Hungarian in five categories . they estimate that the available data sets are 2800 hours across 7500 speakers .
Outcome: The available data sets in spoken Hungarian are compared to other languages and are estimated to be 2800 hours in size . however, their distribution and alignment to real-life tasks are far from optimal indicating the need for larger-scale natural language speech data sets.
VHASR: A Multimodal Speech Recognition System With Vision Hotwords (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models that incorporate audio-related image information do not improve speech recognition performance.
Approach: They propose a novel approach utilizing audio-related image information and set up a multimodal speech recognition system that uses vision as hotwords to enhance the model’s speech recognition capability.
Outcome: The proposed model outperforms unimodal ASR model and achieves SOTA among existing image-based multimodal ASL models.
Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition (2023.acl-long)

Copied to clipboard

Challenge: Existing efforts to improve robustness of audio-visual speech recognition with visual information focus on audio modality . current approaches introduce noise adaptation techniques to improve reliability of AVSR task .
Approach: They propose a visual-invariant modality to strengthen robustness of audio-visual speech recognition (AVSR) it can adapt to any testing noises without dependence on noisy training data, a.k.a., unsupervised noise adaptation.
Outcome: The proposed method outperforms existing state-of-the-arts on visual speech recognition task under various noisy and clean conditions.
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens (2025.findings-acl)

Copied to clipboard

Challenge: Recent Large Language Model (LLM) based AVSR systems incur high computational costs due to high temporal resolution of audio-visual speech.
Approach: They propose an efficient multimodal speech LLM framework that minimizes token length while preserving essential linguistic content.
Outcome: The proposed approach reduces token usage by 86% while using only 3.5 tokens per second.
BLSP-Emo: Towards Empathetic Large Speech-Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: BLSP-Emo model understands both semantics and emotions in speech and generates empathetic responses.
Approach: They propose a language-speech pretraining with emotion support that utilizes existing speech and emotion recognition datasets to create an end-to-end speech-language model.
Outcome: The proposed model can understand both semantics and emotions in speech and generate empathetic responses.
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance (2025.emnlp-main)

Copied to clipboard

Challenge: Autoregressive speech token generation models suffer from hallucinations and undesired vocalizations that do not conform to conditioning inputs.
Approach: They propose an encoder-decoder transformer model that improves contextual adherence of speech token generation LLMs through preference alignment and classifier-free guidance.
Outcome: The proposed model outperforms previous LLM-based models on intelligibility, speaker similarity and naturalness.
On the Robust Approximation of ASR Metrics (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for estimating speech recognition metrics depend on ground truth labels.
Approach: They propose a label-free approach to approximating ASR performance metrics . they embed multimodal embeddings in a unified space for speech and transcription representations .
Outcome: The proposed method outperforms baseline models on speech recognition benchmarks by 50%.
Summarizing Speech: A Comprehensive Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Podcasts and other audiovisual content are becoming more and more a part of everyday communication and the digital age is changing from text to voice.
Approach: They synthesize the current state of the field and highlight the need for realistic evaluation benchmarks and multilingual datasets.
Outcome: The proposed frameworks are based on evaluation protocols and datasets and highlight the need for realistic benchmarks and multilingual datasets.
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification (2024.lrec-main)

Copied to clipboard

Challenge: Similar to humans, animals make extensive use of verbal and non-verbal forms of communication, including audio signals.
Approach: They propose to use self-supervised speech representation models pre-trained on human speech to address dog bark classification tasks.
Outcome: The proposed model improves dog recognition, breed identification, gender classification, and context grounding tasks.
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Currently, there are no publicly available speech recognition datasets in the medical domain due to privacy restrictions.
Approach: They present a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical and 1200h of general-domain speech.
Outcome: The proposed model outperforms state-of-the-art models from 51.8% to 29.6% WER on test set.
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception (2026.acl-long)

Copied to clipboard

Challenge: naively fine-tuning an omni-model on speech recognition and external sound understanding tasks often degrades performance . Xie and Wu's framework, Speech-Hands, recasts the problem as an explicit self-reflection decision.
Approach: They propose a voice-agentic framework that learns one critical omni-understanding skill: trusting itself versus external audio perception.
Outcome: The proposed framework outperforms baseline models on the OpenASR leaderboard by 12.1% WER and high F1 on audio QA decisions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations